Embodied Cognition: A field guide
نویسنده
چکیده
The nature of cognition is being re-considered. Instead of emphasizing formal operations on abstract symbols, the new approach foregrounds the fact that cognition is, rather, a situated activity, and suggests that thinking beings ought therefore be considered first and foremost as acting beings. The essay reviews recent work in Embodied Cognition, provides a concise guide to its principles, attitudes and goals, and identifies the physical grounding project as its central research focus. 2003 Elsevier Science B.V. All rights reserved. For over fifty years in philosophy, and for perhaps fifteen in Artificial Intelligence and related disciplines, there has been a re-thinking of the nature of cognition. Instead of emphasizing formal operations on abstract symbols, this new approach focuses attention on the fact that most real-world thinking occurs in very particular (and often very complex) environments, is employed for very practical ends, and exploits the possibility of interaction with and manipulation of external props. It thereby foregrounds the fact that cognition is a highly embodied or situated activity—emphasis intentionally on all three— and suggests that thinking beings ought therefore be considered first and foremost as acting beings. This shift in focus from Descartes’ “thinking thing”, and the picture of human being and subjectivity it suggests, to a more Heideggerian approach to being in the world, in which agency and interactive coping occupy center stage, is an extremely important development, the implications of which are only just beginning to be fathomed. Very recently a number of books have appeared which detail this shift, and explore in various ways these implications.1 I have selected three of them to discuss in detail here: Cambrian Intelligence by Rodney Brooks [15]; Philosophy in the Flesh by George Lakoff and Mark E-mail address: [email protected] (M.L. Anderson). 1 Books directly related to Embodied Cognition which have appeared just since 1995 include: [3,8,10,13–17, 19,20,22,25,29,35,39–42,55,58,62,65,66,72–74,79,88,91,92,98–100,102,105,111,113,116–120,122,125,126]. 0004-3702/03/$ – see front matter 2003 Elsevier Science B.V. All rights reserved. doi:10.1016/S0004-3702(03)00054-7 ARTICLE IN PRESS S0004-3702(03)00054-7/REV AID:1984 Vol.•••(•••) P.2 (1-40) ELSGMLTM(ARTINT):m1a v 1.155 Prn:29/05/2003; 10:10 aij1984 by:ML p. 2 2 M.L. Anderson / Artificial Intelligence ••• (••••) •••–••• Johnson [72]; and Where the Action Is by Paul Dourish [29]. Together these works span a fair bit of this extremely active and diverse research area, and I will supplement the discussion with references to other works, as appropriate. In, as it were, summarizing the summaries, I hope thereby to provide a concise guide to the principles, attitudes and goals of research in Embodied Cognition. In what follows, I will first very briefly outline the basic foundations and philosophical underpinnings of the classical approach to Artificial Intelligence, against which Embodied Cognition (henceforth EC) is in large part a reaction (Section 1). Section 2 outlines the alternative approach to understanding intelligence pioneered by Rodney Brooks. In that section I also discuss in some detail the central criticisms of classical AI which motivate the situated or embodied alternative, and identify the physical grounding project as the central project of EC. This project is envisioned as a comprehensive version of the symbol grounding problem [52,53]—the well-known question of how abstract symbols can acquire real-world meaning—and centrally involves understanding how cognitive contents (however these are ultimately characterized, symbolically or otherwise) must ultimately ground out in (terms of) the agent’s embodied experience and physical characteristics. In my view, it is the centrality of the physical grounding project that differentiates research in embodied cognition from research in situated cognition, although it is obvious that these two research programs are complementary and closely related. Section 3 outlines the different meanings of embodiment to be found in the EC literature, and discusses each one in relation to the physical grounding project, while Section 4 briefly emphasizes the importance of social situatedness to human level intelligence. In Section 5 I review and evaluate some common criticisms of the EC approach. Finally, Section 6 discusses some implications of this new understanding of intelligence and its grounds for the design of human-computer interfaces. I conclude by briefly recounting the central principles of EC research. 1. Cartesianism, cognitivism, and GOFAI That Descartes is the thinker most responsible for the theoretical duality of mind and body is one of the things that everybody knows; and like most such common knowledge it is not quite accurate. Descartes’ arguments for the separation of body and soul are part of a long legacy of dualistic thinking, in which Plato’s discussion of the immateriality of the soul and the Christian metaphysical tradition which adopted and preserved that discussion played a central role. Indeed, Descartes himself always insisted that, although body and soul were conceptually, and therefore ontologically distinct, they nevertheless formed an empirical unity. What we have inherited from Descartes is a way of thinking about our relation to the world—in particular our epistemological relation to the world— which serves to support and strengthen this ontological stance [91,92]. Descartes’ famous ontological realization that he is a thinking thing is conditioned and tempered by the epistemological admission that all he had accepted as most true had come to him through the senses; yet it is precisely this intrusion of the body between knowledge and world which is in the end unacceptable. The body must be part of the causal order to allow for perceptual interaction, but is therefore both unreliable (a cause of the senses’ deception) and, as it ARTICLE IN PRESS S0004-3702(03)00054-7/REV AID:1984 Vol.•••(•••) P.3 (1-40) ELSGMLTM(ARTINT):m1a v 1.155 Prn:29/05/2003; 10:10 aij1984 by:ML p. 3 M.L. Anderson / Artificial Intelligence ••• (••••) •••–••• 3 were, too reliable (driven by physical forces, and so a potential source of unfreedom).2 Thus, the body is for Cartesian philosophy both necessary and unacceptable, and this ambivalence drives mind and body apart in ways Descartes himself may not have intended. Sensation requires the physicality of the body; the freedom and reliability of human reason and judgment seem to require the autonomy of the soul. We may have no choice about how the world appears to us—we may be to this degree determined by our physical constitution—but we can step back from appearances, and allow reflection and judgment, even if in the end this means adopting the skeptical attitude and simply withholding belief. The postulation of an inner arena of disengaged representation, influenced by experience but governed by reasons and not causes, is surely a natural way to try to account for this possibility.3 These tensions clearly play a role in another aspect of the Cartesian world-view, not much emphasized in these contexts, but perhaps of greater importance: the discontinuity between humans and animals. For Descartes, animals are mere mechanisms, complex and interesting to be sure, but physical automata nevertheless [23, Vol. III, pp. 365–366, 374]. They do have sensation, because all that is needed for that is the proper organs.4 However, they lack thought, perhaps the most important evidence for which is the fact that they lack language. This denial that sensing and acting in the world require thinking, and the concomitant identification of thinking with the higher-order reasoning and abstraction paradigmatically displayed in language use is perhaps the true heart of the Cartesian attitude. Indeed, I believe that it is primarily from this inheritance that the central attitudes and approach of cognitivism can be derived. Simply put, cognitivism is the hypothesis that the central functions of mind—of thinking—can be accounted for in terms of the manipulation of symbols according to explicit rules. Cognitivism has, in turn, three elements of note: representation, formalism, and rule-based transformation. First and foremost is the idea that cognition centrally involves representation; cognitivism is committed to the existence of “distinct, identifiable, inner states or processes”—that is, the symbols—“whose systemic or functional role is to stand in for specific features or states of affairs” [20, p. 43]. However, just as is the case in modern logic, it is the form of the symbol (or the proposition of which the symbol is a part) and not its meaning that is the basis of its rule-based transformation. To some degree, of course, such formal abstraction is a necessary condition for representation—the token for ‘green’ in my mental lexicon is not itself green, nor does it necessarily share any of the other properties of green. Indeed, the relation between sign and signifier seems in this sense necessarily arbitrary,5 and this thereby enforces a kind of distance between the inner 2 The thesis of physical determinism as we have come to formulate it awaited later developments in mechanics and mathematics, and attained its first recognizably contemporary form in the work of Pierre-Simon Laplace (1749–1827); still, mechanism was an important part of Descartes’ world-view. See, e.g., [28]. 3 See [49] for a very nice account and critique of Cartesian disengagement, from a Heideggerian perspective. 4 “Please note that I am speaking of thought, and not of life or sensation. I do not deny life to animals, since I regard it as consisting simply in the heat of the heart; and I do not even deny sensation, in so far as it depends upon a bodily organ. Thus my opinion is not so much cruel to animals as indulgent to human beings—at least to those who are not given to the superstitions of Pythagoras—since it absolves them from the suspicion of crime when they eat or kill animals” [23, Vol. III, p. 366]. 5 But see [107] for some suggestions to the contrary. ARTICLE IN PRESS S0004-3702(03)00054-7/REV AID:1984 Vol.•••(•••) P.4 (1-40) ELSGMLTM(ARTINT):m1a v 1.155 Prn:29/05/2003; 10:10 aij1984 by:ML p. 4 4 M.L. Anderson / Artificial Intelligence ••• (••••) •••–••• arena of symbol processing and the external world of meaning and action. Still, as we will see in more detail later, this formal abstraction is nevertheless a matter of degree, and this aspect of cognitivism is separate—and separately criticizable—from the general issue of representation. This brings us to third important aspect of cognitivism: the commitment to explicitly specifiable rules of thought. This commitment follows naturally from the others, for having disconnected the form of a symbol from its meaning, cognitivism rules out the possibility of content-sensitive processing, and so requires formal rules to govern the transformation from one cognitive state to another.6 This rather too simple account of cognitivism as the commitment to an account of cognition based on the manipulation of abstract representations by explicit formal rules, brings us to GOFAI (Good Old Fashioned Artificial Intelligence), which has been since its inception, and perhaps still is, the dominant research paradigm in Artificial Intelligence. GOFAI might be loosely defined as any attempt to produce machine intelligence by methods which primarily reflect Cartesian and cognitivist attitudes. The looseness of this definition is meant to emphasize the fact that there are probably no research projects in AI which are purely GOFAI, despite the caricatures which abound in the critical literature. Nonetheless, that many research projects, particularly those from AI’s foundational period,7 display the influences of Cartesian and cognitivist thinking is undeniable. A particular example will help illustrate these influences, and serve as a reference point for later discussion. Perhaps the most impressive current effort along these lines is Cyc (as in enCYClopedia), the now almost 20-year-old project to create a general-purpose common sense reasoner. The approach is straightforwardly cognitivist: The Cyc knowledge base (KB) is a formalized representation of a vast quantity of fundamental human knowledge: facts, rules of thumb, and heuristics for reasoning about the objects and events of everyday life. The medium of representation is the formal language CycL. . . The KB consists of terms—which constitute the vocabulary of CycL—and assertions which relate those terms [24]. It is the conviction of Cyc’s designers that the result of learning, and the precondition for more (and more sophisticated) learning, is the acquisition and retention of many, many, many facts about the world and how it works. The most basic knowledge is the sort of commonsense stuff that we take utterly for granted: Consider Dr. Dave Bowman, mission commander of Discovery. His experiences as an engineering or astrophysics student and his astronaut training qualified him to lead the mission. Before that, his high school and undergraduate college prepared him for graduate school, and before that, he learned in elementary and middle school the 6 In point of fact, the question of how to differentiate symbolic from non-symbolic, rule-based from nonrule-based cognitive systems—how to determine, for instance, whether connectionist implementations of simple reasoning employ representations manipulated by rules—is extremely vexed. See [9,60,123] for a more thorough discussion of these issues. 7 Dreyfus [32] suggests 1957–1977; Brooks extends this range to about 1981, and suggests further that up until 1991 “there was very little change in the underlying assumptions about the models of thought” [15, p. 153]. ARTICLE IN PRESS S0004-3702(03)00054-7/REV AID:1984 Vol.•••(•••) P.5 (1-40) ELSGMLTM(ARTINT):m1a v 1.155 Prn:29/05/2003; 10:10 aij1984 by:ML p. 5 M.L. Anderson / Artificial Intelligence ••• (••••) •••–••• 5 fundamentals he needed for high school. And long before—here we get to some very important stuff indeed—his early experiences as a baby and toddler prepared him for kindergarten and first grade. He learned, for instance, to talk and that people generally prepare food in kitchens. He learned that if you leave something somewhere, it often remains there, at least for a while. He learned that chairs are for sitting on and that pouring milk from one glass to another differently shaped glass doesn’t change the total volume of the milk. He learned that there’s no air in outer space so he’d better not forget the helmet of his spacesuit if he’s going walking in space. He learned all this—and a million other things [76, pp. 195–196]. For the purposes of this essay, we might call the supposition that we learn these sorts of things: if something is not supported it falls, that learning entails representing that knowledge in this sort of way: ist((∀x¬supported(x) → falls(x)),NTP)8, and that reasoning is the manipulation of such representations according to formal, logic-like rules, the central hypothesis of GOFAI. It is against this hypothesis that research in EC is principally arrayed. 2. Cambrian intelligence Although Hubert Dreyfus must be given a great deal of credit for first drawing attention to the limitations of GOFAI [30], I think it can be argued that the single figure most responsible for the new AI is Rodney Brooks. A collection of his seminal papers has recently appeared [15] which together comprise not just a sustained critique of the central hypothesis of GOFAI, but the outlines of a philosophically astute and scientifically principled alternative. The book is split into two sections, Technology and Philosophy, with the former containing such papers as “A Robust Layered Control System for a Mobile Robot”, and the latter “Intelligence without Representation” and “Intelligence without Reason”, among others. “Intelligence without Reason”, by far the longest paper in the collection (pp. 133–186), is also the most complete statement of Brooks’ guiding philosophy, and today stands as the overall best introduction to the principles and motivations of the new AI. As we have seen, traditional AI is characterized by an understanding of intelligence which foregrounds the notions of thought and reason, and adopts certain conventions for approaching these which centrally involve the creation of representations, and the deployment of high-level cognitive skills such as planning and problem solving. For Brooks, however, such an approach “cannot account for large aspects of what goes into intelligence” (p. 134). In contrast to this high-level or top-down approach to intelligence, Brooks advocates studying intelligence from the bottom up, and specifically urges us to recall our evolutionary lineage. As evolved creatures, human beings are largely continuous with our forebears, and we have inherited from them a substrate of capacities and systems 8 ist(x, y) means x is true in theory y: in this case y = NTP, the Naive Theory of Physics, one the many microtheories employed in Cyc [48]. ARTICLE IN PRESS S0004-3702(03)00054-7/REV AID:1984 Vol.•••(•••) P.6 (1-40) ELSGMLTM(ARTINT):m1a v 1.155 Prn:29/05/2003; 10:10 aij1984 by:ML p. 6 6 M.L. Anderson / Artificial Intelligence ••• (••••) •••–••• for meeting our needs in, and generally coping with a given environment. From such considerations follows the perhaps reasonable, but decidedly un-Cartesian thought: “The study of that substrate may well provide constraints on how higher level thought in humans could be organized” (p. 135, emphasis in original). As we will see, this tendency to emphasize, on evolutionary grounds, the continuity between humans and other animals, and the converse willingness to see in animals instances of intelligent behavior, is an extremely important motivation for the study of EC. But while such thinking may provide justification for pursuing a certain sort of research into cognition, it does not by itself constitute a critique of other approaches. So what, exactly, is wrong with GOFAI? This, of course, is a difficult and vexed question, because arguments over theoretical principles, such as whether humans really represent things in terms of proposition-like mental entities, have so far proved indecisive, there being strong theoretical support on both sides of the issue; so far, neither sufficiently compelling empirical evidence, nor widespread agreement on decisive foundational assumptions has been achieved.9 Similarly, claims about the deficiencies of particular existing systems can always be deflected, in most cases legitimately, by pointing out that any engineering project is governed by real resource limitations—temporal limitations most significantly, including the need for the graduate students who primarily build such systems to graduate—and that the next iteration will address the noted shortcomings. Still, even without the sort of definitive experiment or event which occasionally marks a scientific revolution [70], it should be possible to look at the trajectory and tendencies of any given research program, and come to some sort of tentative assessment of its promise. Thus does Brooks approach work in AI and robotics. Surveying some early work, e.g., Skakey [87], CART [84] and Hilare [47], he comments: All of these systems used offboard computers . . . and operated in mostly static environments. All of these robots operated in environments that at least to some degree had been specially engineered for them. They all sensed the world and tried to build two or three dimensional world models of it. Then, in each case, a planner could ignore the actual world, and operate in the model to produce a plan of action for the robot to achieve whatever goal it had been given. . . . We will call this framework the sensemodel-plan-act framework, or SMPA for short (pp. 136–137). One of the problems that Brooks identifies with this approach is that it is insufficiently dynamic. Once the system has built its model, it will not notice changes in its actual environment. Then, when it comes time to act (and given the slowness of the systems, this time may be long in coming) the world may well have changed in ways which make the plan obsolete. Depending on the sensitivity of the system, this will either result in the re-initiation of the SMPA cycle, or the execution of an ineffective plan. Neither result is optimal, to say the least. Of course, one way to address this problem is to improve 9 Indeed, as far as agreement or progress on foundational theoretical issues, there has been little if any progress in the last 10 years. See [68], which presents a set of controversies which must be counted as yet controversial. It is perhaps worth noting, however, that these issues are being discussed; it was Kirsh’s impression in 1991 that they had been largely ignored. ARTICLE IN PRESS S0004-3702(03)00054-7/REV AID:1984 Vol.•••(•••) P.7 (1-40) ELSGMLTM(ARTINT):m1a v 1.155 Prn:29/05/2003; 10:10 aij1984 by:ML p. 7 M.L. Anderson / Artificial Intelligence ••• (••••) •••–••• 7 performance, especially in the extremely expensive process of perception and model construction. But for Brooks and his followers, this is to miss the point that the SMPA framework is by its nature too expensive, and therefore biologically implausible. As an alternative they encourage the adoption of highly reactive models of perception, bypassing the representational level altogether. As he has famously formulated it: We have reached an unexpected conclusion (C) and have a rather radical hypothesis (H). (C) When we examine very simple level intelligence we find that explicit representations and models of the world simply get in the way. In turns out to be better to use the world as its own model. (H) Representation is the wrong unit of abstraction in building the bulkiest parts of intelligent systems (pp. 80–81). It is true that despite vast improvements in the speed of microprocessors, and significant advances in such areas as computer vision, knowledge representation, non-monotonic reasoning, and planning there has yet to be an SMPA system that can operate in a complex, real-world environment on biologically realistic time-scales. The twin scale-up of environmental richness and real-time dynamics has so far proved insurmountable. On the other side of the coin, all these areas are advancing, and so perhaps the achievement of real-time SMPA intelligence is just a matter of time. Showing that something has not yet happened is a long way from showing it won’t. Yet there are reasons to suspect the latter. Consider first the central problem of planning, that is, of figuring out what to do next, given, say, a certain goal, and a current situation. On the SMPA model, this process involves examining the world model, and mapping on to that model a series of steps which will achieve the desired state. This abstract description hides two kinds of complication, which I will call dynamics and relevance. The problem of dynamics we have mentioned already; if the world changes, then the plan will have to be likewise adjusted to fit the new situation. Within the SMPA framework, there are two approaches to this problem. The first is to include the dynamics of the world in the model, and to plan in terms of the expected changes in the environment. Naturally, this only pushes the problem back one step, for now we have to monitor whether or not the changes are the expected ones, and re-plan when they are not. This suggests a second approach, to include in the initial plans contingent sub-plans to be executed in the case of the various possibilities for the future states of the world. I suspect it is obvious both that this would result in an unmanageable computational explosion, and that real-world planning is nothing like this. As difficult as the problem of dynamics is, there is yet another problem implicit in the above discussion. For one doesn’t want to re-plan in the face of every change, only those which are relevant, that is, which are likely to affect the achievability of the goal. Thus, for a heavy robot moving across a room the location and dynamics of big, solid objects is likely relevant, but the speed and direction of the draft from the open window is not. Unless the task is to carry a stack of papers. Likewise, the broken and deeply pitted floor tiles make no difference to a robot with big spongy wheels, but might matter to one with different means for locomotion. In general, what counts as a relevant fact worth noticing— ARTICLE IN PRESS S0004-3702(03)00054-7/REV AID:1984 Vol.•••(•••) P.8 (1-40) ELSGMLTM(ARTINT):m1a v 1.155 Prn:29/05/2003; 10:10 aij1984 by:ML p. 8 8 M.L. Anderson / Artificial Intelligence ••• (••••) •••–••• say, whether something falls in the general class of obstacle or not—will depend both on the capacities of the agent and the task to be performed. This is obvious enough, to be sure, but turns out to be notoriously difficult to implement in a representational system of any size. For let us suppose our agent to have modeled a sufficiently complex environment; we would expect this to require many thousands of ‘facts’—let’s say a large number equal to m. According to the above ‘obvious’ requirement for agency, our agent would need to know, for each such fact, whether or not it was relevant to the proposed action. And to know this requires that the agent know a good deal about each of the actions in its repertoire (more facts), and further know how to tell whether or not a given (environmental) fact was relevant to a given behavioral fact. Given a number of such behavioral facts n, and assuming a straightforward set of relevance rules p could be written, determining the relevance of the various aspects of the environment to any one action will minimally require m ∗ n ∗ p comparisons. Further, assuming that the robot is reasoning about what it knows (not to mention that the environment is itself changing, and perhaps the robot is noticing this), the number m is constantly increasing, as new facts are added to the KB through reasoning and observation. It is this sort of performance problem that is behind Dennett’s spoof of the third-generation, newly re-designed robot trying to figure out how to remove its battery from the room it shares with a time-bomb: Back to the drawing board. ‘We must teach it the difference between relevant implications and irrelevant implications’, said the designers, ‘and teach it to ignore the irrelevant ones’. So they developed a method of tagging implications as either relevant or irrelevant to the project at hand, and installed the method in their next model, the robot-relevant-deducer, or R2D1 for short. When they subjected R2D1 to the test that had so unequivocally selected its ancestors for extinction, they were surprised to see it sitting, Hamlet-like, outside the room containing the ticking bomb, the native hue of its resolution sicklied o’er with the pale cast of thought, as Shakespeare (and more recently Fodor) has aptly put it. ‘Do something!’ they yelled at it. ‘I am’, it retorted. ‘I’m busily ignoring some thousands of implications I have determined to be irrelevant. Just as soon as I find an irrelevant implication, I put it on the list of those I must ignore, and. . .’ the bomb went off [27, p. 129]. It may be that the relevance problem is merely practical, but if so, it is a very deep practical problem that borders on the theoretical. At the very least it is a fundamental issue for any representational system, for at the root of the relevance problem is that most basic question for any representation: what should be modeled, and what ignored or abstracted away? One moral which might be derived from the problem of relevance—a weak form of the more radical moral that Brooks actually draws—is that it doesn’t make sense to think about representing at all unless one knows what one is representing for. The SMPA model envisions the infeasible task of deriving relevance from an absolute world model; but if we step back a pace we would realize that no system can ever have such an absolute (or, context-free) world model in the first place. Every representer is necessarily selective, and a good representation is thereby oriented toward a particular (sort of) use by a particular (sort of) agent. Just as the problem of dynamics seems to suggest the inevitability of short, ARTICLE IN PRESS S0004-3702(03)00054-7/REV AID:1984 Vol.•••(•••) P.9 (1-40) ELSGMLTM(ARTINT):m1a v 1.155 Prn:29/05/2003; 10:10 aij1984 by:ML p. 9 M.L. Anderson / Artificial Intelligence ••• (••••) •••–••• 9 incomplete plans, so the problem of relevance pushes us towards adopting more limited representations, closely tied to the particularities, and therefore appropriate to the needs, of a given agent. For Brooks, the failure of the SMPA model indicates the nature of goal-orientation (which is, after all, the cognitive phenomenon that planning is intended to capture) should be re-thought. The above problems seem to suggest their own solution: shorter plans, more frequent attention to the environment, and selective representation. But the logical end of shortening plan length is the plan-less, immediate action; likewise the limit of more frequent attention to the environment is constant attention, which is just to use the world as its own model. Finally, extending the notion of selective representation leads to closing the gap between perception and action, perhaps even casting perception largely in terms of action.10 Thus, the problems of dynamics and relevance push us toward adopting a more reactive, agent-relative model of real-world action, what I call situated goal-orientation.11 Brooks writes, Essentially the idea is to set up appropriate, well conditioned, tight feedback loops between sensing and action, with the external world as the medium for the loop. . . . We need to move away from state as the primary abstraction for thinking about the world. Rather, we should think about processes which implement the right thing. We arrange for certain processes to be pre-disposed to be active and then given the right physical circumstances the goal will be achieved. The behaviors are gated on sensory inputs and so are only active under circumstances where they might be appropriate. Of course, one needs to have a number of pre-disposed behaviors to provide robustness when the primary behavior fails due to being gated out. As we keep finding out what sort of processes implement the right thing, we continually redefine what the planner is expected to do. Eventually, we won’t need one (p. 109). One is certainly entitled to doubt whether this sort of selective-reactive attunement to the environment can account for the complex tasks routinely faced by the typical soccer mom, as she shuttles her various children to their various activities at the right time, in the right order, meanwhile figuring out ways to work in laundry, dry cleaning and grocery shopping [4]. Likewise, dialog partners seem to engage in complex 10 This idea, which I think is ultimately very powerful, suggests a notion familiar from phenomenology that the perceptual field is always already an action-field—that the perceived world is always known in terms directly related to an agent’s current possibilities for future action. One way of cashing this out is in terms of affordances, the perceived availability of things to certain interventions [46], so that the world, as it were, constantly invites action; another is Brooks’ notion of a specialized repertoire of behaviors selectively gated by environmental stimuli, which is, in turn, not unrelated to Minsky’s notion of the society of mind [83]. More generally, these notions are related to the theory of direct perception, which is usefully and thoroughly discussed in [1] and [19] Chapters 11–12. 11 It should be noted that the relevant neuroscientific studies suggest evidence for both agentor body-relative representations of action [12,63] and objective, or world-relative representations [44,45]. As with many such tensions, the path of understanding is not likely to be either/or, but rather both/and, with the hardest work to be done in understanding the relation between the two kinds of representations. ARTICLE IN PRESS S0004-3702(03)00054-7/REV AID:1984 Vol.•••(•••) P.10 (1-40) ELSGMLTM(ARTINT):m1a v 1.155 Prn:29/05/2003; 10:10 aij1984 by:ML p. 10 10 M.L. Anderson / Artificial Intelligence ••• (••••) •••–••• behaviors—remembering what has been said; deciding what to say next in light of overall conversational goals; correcting misapprehensions, and reinterpreting past interactions in light of the corrections, thereby (sometimes) significantly altering the current dialog state—which defy analysis without representations.12 Indeed, it is a vice too often indulged by scientists working in EC to make the absence of representations a touchstone of virtue in design, and to therefore suppose that, just as do the creatures they devise in the lab, so too must humans display an intelligence without representations. Representationphobia is a distracting and ultimately inessential rhetorical flourish plastered over a deep and powerful argument. For rather than targeting representations per se, the central argument of EC instead strikes at their nature and foundation, thereby forcefully, and to my mind successfully raising the question of whether GOFAI has chosen the right elements with which—the right foundation on which—to build intelligence. For suppose that robotic soccer mom does need representations, does need an abstract, symbolic planner. Still, if EC is on anything like the right track she cannot live by symbols alone; her representations must be highly selective, related to her eventual purposes, and physically grounded. This strongly suggests that her faculty of representation should be linked to, and constrained by, the ‘lower’ faculties which govern such things as moving and acting in a dynamic environment,13 without questioning the assertion that complex agency requires both reactive and deliberative faculties.14 The central moral coming from EC is not that traditional AI ought to be given up, but rather that in order to incorporate into real-world agents the sort of reasoning which works so well in expert systems, ways must be found to systematically relate the symbols and rules of abstract reasoning to the more evolutionarily primitive mechanisms which control perception and action. As suggested, this may well involve rethinking the nature and bases of representation, as well as a willingness constrain the inner resources available to automated reasoners (at least those intended for autonomous agents) to items which can be connected meaningfully to that agent’s perceptions and actions (either directly or through connected resources themselves appropriately related to perception and action). This returns us to Brooks’ evolutionary musings, and a passage apparently so instructive that it appears twice: It is instructive to reflect on the way in which earth-based biological evolution spent its time. Single-cell entities arose out of the primordial soup roughly 3.5 billion years ago. A billion years passed before photosynthetic plants appeared. After almost another billion and a half years, around 550 million years ago, the first fish and vertebrates 12 [123] discusses the general importance of planning to cognitive agents, as part of a critique of research in situated action. 13 There has been a good deal of interesting work on the relation of higher faculties to lower, from an evolutionary perspective. A good place to start is [108]; see also [88,125]. 14 Research on cognitive agents [43,50,61,77,115] explicitly begins with this both/and attitude. My own work in philosophy and AI is aimed primarily at understanding how deliberative, representing agents can cope with a complex, changing environment. Thus I have been working with Don Perlis on time-situated non-monotonic reasoning, and especially on the problem of uncertainty [4,11,18], with Tim Oates on automated reasoning with grounded symbols [90], and with Gregg Rosenberg on a specification for embodied, action-oriented concepts and representations [106]. ARTICLE IN PRESS S0004-3702(03)00054-7/REV AID:1984 Vol.•••(•••) P.11 (1-40) ELSGMLTM(ARTINT):m1a v 1.155 Prn:29/05/2003; 10:10 aij1984 by:ML p. 11 M.L. Anderson / Artificial Intelligence ••• (••••) •••–••• 11 appeared, and then insects 450 million years ago. Then things started moving fast. Reptiles arrived 370 million years ago, followed by dinosaurs at 330 and mammals at 250 million years ago. The first primates appeared 120 million years ago and the immediate predecessors to the great apes a mere 18 million years ago. Man arrived in roughly his present form 2.5 million years ago. He invented agriculture a mere 10,000 years ago, writing less than 5000 years ago and ‘expert’ knowledge only over the last few hundred years. This suggests that problem solving behavior, language, expert knowledge and application, and reason, are all pretty simple once the essence of being and reacting are available. That essence is the ability to move around in a dynamic environment, sensing the surroundings to a degree sufficient to achieve the necessary maintenance of life and reproduction. This part of intelligence is where evolution has concentrated its time–it is much harder (p. 81 also pp. 115–116). This is the physically grounded part of animal systems . . . these groundings provide the constraints on symbols necessary from them to be truly useful (p. 116). This is perhaps the most important theoretical upshot of Brooks’ work: the knowledgeis-everything approach of GOFAI, exemplified especially by Douglas Lenat and the others involved in Cyc, can never work, for such a system lacks the resources to ground its representations [52,53]. More importantly, the structure and type of representation employed in Cyc, being unconstrained by the requirements any grounded, agency-oriented components, may well turn out to be unsuited for direct grounding. Despite the “implicit assumption that someday the inputs and outputs will be connected to something which will make use of them” (p. 155) the system in fact relies entirely on human interpretation to give meaning to its symbols, and therefore implicitly requires the still mysterious—and no doubt highly complex—human grounding capacity to serve as intermediary between its outputs and real-world activity. This is one reason why EC researchers tend to prefer a bottom-up approach to grounding, and often insist on working only with symbols and modes of reasoning which can be straightforwardly related to perception and action in a particular system. A system like Cyc which relies on human faculties to give meaning to its symbols may well require a fully implemented human symbol grounder, including human-level perceptual and behavioral capacities, to be usefully connected to the real world. Grounding the symbol for ‘chair’, for instance, involves both the reliable detection of chairs, and also the appropriate reactions to them.15 These are not unrelated; ‘chair’ is not a concept definable in terms of a set of objective features, but denotes a certain kind of thing for sitting. Thus is it possible for someone to ask, presenting a tree stump in a favorite part of the woods, “Do you like my reading chair?” and be understood. An agent who has grounded the concept ‘chair’ can see that the stump is a thing for sitting, and is therefore (despite the dearth of objective similarities to the barcalounger in the living room, 15 Consider, in this regard, the difference between the perceptual judgments “There’s a fire here” and “There’s a fire there”, or “There’s a dollar here”. Surely appropriate grounding has not been achieved by a system which does not react differently in each case. ARTICLE IN PRESS S0004-3702(03)00054-7/REV AID:1984 Vol.•••(•••) P.12 (1-40) ELSGMLTM(ARTINT):m1a v 1.155 Prn:29/05/2003; 10:10 aij1984 by:ML p. 12 12 M.L. Anderson / Artificial Intelligence ••• (••••) •••–••• and despite also being a tree stump) a chair.16 Simply having stored the fact that a chair is for sitting is surely not sufficient ground for this latter capacity. The agent must know what sitting is and be able to systematically relate that knowledge to the perceived scene, and thereby see what things (even if non-standardly) afford sitting. In the normal course of things, such knowledge is gained by mastering the skill of sitting (not to mention the related skills of walking, standing up, and moving between sitting and standing), including refining one’s perceptual judgments as to what objects invite or allow these behaviors; grounding ‘chair’, that is to say, involves a very specific set of physical skills and experiences. A further problem is the holistic nature of human language and reasoning: ‘chair’ is closely related to other concepts like ‘table’. One is entitled to wonder whether knowing what sort of chairs belong at a table is part of the mastery of ‘table’ or ‘chair’. It is unlikely that clear boundaries can be drawn here; knowing one partly involves knowing the other. Likewise, ‘chair’ is related to ‘throne’, so that it is not clear whether we should say of someone who walked up and sat in the King’s throne that she failed to understand what a throne was, or failed to understand what a chair was (and wasn’t). Given that these concepts are semantically related, that there is a rational path from sentences with ‘chair’ to sentences with ‘table’ or ‘throne’, any agent who hopes to think with ‘chair’ had better have grounded ‘table’ and ‘throne’, too. The point is not that Cyc’s symbols can never be grounded or given real-world meaning. Rather, the basic problem is that the semantic structures and rules of reasoning that Cyc uses are rooted in an understanding of what these concepts mean for us, and therefore they can be usefully (and fully) grounded only for a system which is likewise very much like us. This includes not just basic physical and perceptual skills, but also (as the throne example demonstrates) being attuned to the social significance of objects and situations, for seeing the behaviors that an object affords may depend also on knowing who one is, and where one fits not just in the physical, but also the social world. An android with this sort of mastery is surely not in our near future, but a lesser robot using Cyc for its knowledge base is just wasting clock cycles, and risks coming to conclusions that it cannot itself interpret or put to use—or worse, which it will misinterpret to its detriment. For EC, the only useful reasoning is grounded reasoning.17 The notion that grounding is at the root of intelligence, and is the place to look for the elusive solution to the relevance problem—for grounding provides the all-important 16 Lying behind this claim is the theory of direct, or ecological, perception [46], which for reasons of space has not been discussed in detail here. In short, the idea is (1) at least some perception does not involve inference from perception to conception, but is rather direct and immediate, and (2) perception is generally action-oriented, and the deliverances of the perceptual system are often cast in action-relative, and even, as in the case of affordances, action-inviting terms. In illustration of claim (1), Clancey [19] suggests that one’s identification, while blindfolded, of the object in one’s hand as an apple is unmediated and direct; it is part of one’s immediate perceptual experience that the object is an apple. He contrasts this with the case where one counts the bumps on the bottom of the apple to infer that it is a Red Delicious. The latter involves inference from perceptually gathered information; the former does not (pp. 272–273). Claim (2) we have already discussed, in, e.g., note 10, and immediately above. The theory of direct perception is usefully and thoroughly discussed in [1] and [19] Chapters 11–12. 17 Note that this doesn’t mean that Cyc is not a useful tool. The point is that Cyc is only useful for an agent that can ground its symbols. Humans are currently the only such agents known. ARTICLE IN PRESS S0004-3702(03)00054-7/REV AID:1984 Vol.•••(•••) P.13 (1-40) ELSGMLTM(ARTINT):m1a v 1.155 Prn:29/05/2003; 10:10 aij1984 by:ML p. 13 M.L. Anderson / Artificial Intelligence ••• (••••) •••–••• 13 constraints on representation and inference with which the purely symbolic approach has such trouble—Brooks calls “the physical grounding hypothesis” (p. 112). [Cyc employs] a totally unsituated, and totally disembodied approach . . . [110] provides a commentary on this approach, and points out how the early years of the project have been devoted to finding a more primitive level of knowledge than was previously envisioned for grounding higher levels of knowledge. It is my opinion, and also Smith’s, that there is a fundamental problem still, and one can expect continued regress until the system has some form of embodiment (p. 155). As interesting and revolutionary as Brooks’ contribution to AI and robotics has been, the ultimate success of the research program depends on whether and how the physical grounding hypothesis can be cashed out. This will mean, among other things, coming to a concrete understanding of phrases like “some form of embodiment”, and specifying in detail how embodiment in fact acts to constrain representational schema and influence higher-order cognition. As I have already mentioned, I take this concern with physical grounding to be a central defining characteristic of research in EC, differentiating it from research in situated cognition. Thus it is to recent contributions to this project—which I will call the physical grounding project—that we now turn. 3. Embodiment and grounding It is, of course, obvious that embodiment is of central interest to EC. Yet it is also clear that the field has yet to settle on a shared account of embodiment: indeed, there is little explicit discussion of its meaning, as if it were a simple term in little need of analysis. When considered in the abstract, primarily as an indication of the area of focus in contrast to that indicated by an interest in ‘symbols’, this is no doubt good enough. But having thus focused attention on embodiment, it is incumbent on the field to say something substantial about its meaning. In this project, the resources of phenomenology should not be overlooked, especially [56,81,82], despite their difficulty.18 Naturally, a single section of a field review is not the place to try to give the required account. Instead of trying to provide a unified theory, I will lay out the general terrain, and say something about the various meanings of embodiment and their relation to the physical grounding project. Perhaps the most basic, but also the most theoretical, abstract way to explain what it means to focus attention on the embodiment of the thinking subject comes from MerleauPonty and Heidegger. Heidegger, for instance, shows—especially in his celebrated analysis of being-in-theworld—that the condition of our forming disengaged representations of reality is that 18 [31,59,121] are all, in part, attempts to capture the insights of phenomenology in a way directly useful to cognitive science. [117] is a more strictly philosophical attempt to give a thorough account of embodiment and its relation to mind. [51] relates representations to intervention in the context of scientific practice. More work along these various lines is needed. ARTICLE IN PRESS S0004-3702(03)00054-7/REV AID:1984 Vol.•••(•••) P.14 (1-40) ELSGMLTM(ARTINT):m1a v 1.155 Prn:29/05/2003; 10:10 aij1984 by:ML p. 14 14 M.L. Anderson / Artificial Intelligence ••• (••••) •••–••• we be already engaged in coping with our world, dealing with the things in it, at grips with them. . . . It becomes evident that even in our theoretical stance we are agents. Even to find out about the world and formulate disinterested pictures, we have to come to grips with it, experiment, set ourselves to observe, control conditions. But in all this, which forms the indispensable basis of theory, we are engaged as agents coping with things. It is clear that we couldn’t form disinterested representations any other way. What you get underlying our representations of the world—the kinds of things we formulate, for instance, in declarative sentences—is not further representations but rather a certain grasp of the world that we have as agents in it [114, pp. 432–433]. Likewise, Merleau-Ponty argues that perception and representation always occur in the context of, and are therefore structured by, the embodied agent in the course of its ongoing purposeful engagement with the world. Representations are therefore ‘sublimations’ of bodily experience, possessed of content already, and not given content or form by an autonomous mind; and the employment of such representations “is controlled by the acting body itself, by an ‘I can’ and not an ‘I think that’ ” ([59, pp. 108–109], see also [33]). A full explanation of the significance of this claim would require an excursion into the bases of the categories of representation and the transcendental unity of apperception in Descartes and Kant, for which there is no room here (but see [117]). However, the most immediate point is fairly straightforward: the content and relations of concepts—that is, the structure of our conceptual schema—is primarily determined by practical criteria, rather than abstract or logical ones. Likewise, experience, which after all consists of ongoing inputs from many different sources, is unified into a single object of consciousness by, and in terms of, our practical orientation to the world: “the subject which controls the integration or synthesis of the contents of experience is not a detached spectator consciousness, an ‘I think that’, but rather the body-subject in its ongoing active engagement with [the world]” [59, p. 111]. Considering the problem at this level of abstraction is an important ongoing philosophical project, for re-interpreting the human being in terms which put agency rather than contemplation at its center will have eventual implications for many areas of human endeavor.19 Continued engagement with phenomenology and related areas is essential to this re-understanding. However, for the more narrow project of coming to grips with the details of physical grounding, it is perhaps better to turn our attention elsewhere. One extremely important contribution to the physical grounding project has been the ongoing work of George Lakoff and Mark Johnson. Since at least 1980, with the publication of Metaphors We Live By [71], they have been arguing that the various domains of our mental life are related to one another by cross-domain mappings, in which a target domain inherits the inferential structure of the source domain. For instance, the concept of an argument maps on to the concept of war (Argument is War) and therefore reasoning about arguments naturally follows the same paths. It is important to see that we don’t just talk about arguments in terms of war. We can actually win or lose arguments. We see the person we are arguing with as an opponent. 19 Not the least important of which is Ethics. See, e.g., [122]. ARTICLE IN PRESS S0004-3702(03)00054-7/REV AID:1984 Vol.•••(•••) P.15 (1-40) ELSGMLTM(ARTINT):m1a v 1.155 Prn:29/05/2003; 10:10 aij1984 by:ML p. 15 M.L. Anderson / Artificial Intelligence ••• (••••) •••–••• 15 We attack his positions and defend our own. We gain and lose ground. We plan and use strategies. If we find a position indefensible, we can abandon it and take a new line of attack. . . . It is in this sense that the Argument is War metaphor is one that we live by in this culture; it structures the actions we perform in arguing [71, p. 4]. Lakoff and Johnson’s most recent contribution along these lines is Philosophy in the Flesh [72], which is the strongest, most complete statement to date of the claim that all cognition eventually grounds out in embodiment. Philosophy in the Flesh is in many ways a flawed work, for reasons stemming primarily from the authors’ apparent disdain for the large amount of related work in philosophy, artificial intelligence, and cognitive science. I have briefly detailed some of these limitations in [93]. However, because of the importance and promise of its central argument and approach, I will focus here on the positive contribution of the book. As an illustration of how a given example of higher-order cognition can be traced back to its bodily bases, consider the metaphorical mapping “Purposes are Destinations”, and the sort of reasoning about purposes which this mapping is said to encourage. We imagine a goal as being at some place ahead of us, and employ strategies for attaining it analogous to those we might use on a journey to a place. We plan a route, imagine obstacles, and set landmarks to track our progress. In this way, our thinking about purposes (and about time, and states, and change, and many other things besides) is rooted in our thinking about space. It should come as no surprise to anyone that our concepts of space—up, down, forward, back, on, in—are deeply tied to our bodily orientation to, and our physical movement in, the world. According to Lakoff and Johnson, every domain which maps onto these basic spatial concepts (think of an upright person, the head of an organization, facing the future, being on top of things) thereby inherits a kind of reasoning—a sense of how concepts connect and flow—which has its origin in, and retains the structure of, our bodily coping with space. Like most work in EC, Philosophy in the Flesh has little explicitly to say about what the body is. Nevertheless, it is possible to identify four aspects of embodiment that each play a role in helping shape, limit and ground advanced cognition: physiology; evolutionary history; practical activity; and socio-cultural situatedness. We will discuss each in turn.
منابع مشابه
Embodied artificial intelligence
Mike Anderson1 has given us a thoughtful and useful field guide: Not in the genre of a bird-watcher’s guide which is carried in the field and which contains detailed descriptions of possible sightings, but in the sense of a guide to a field (in this case embodied cognition) which aims to identify that field’s general principles and properties. I’d like to make some comments that will hopefully ...
متن کاملThe Explanation of effectiveness of student's lived experience in the architectural training process
Abstract: Architecture as a built environment has an important role in the quality of experience. Architecture experience is one of the educational strategies for knowing architecture. In some educational approaches, experience means observation, while based on the embodied cognition, environment experience is more than an observation. According to this approach, lived experience is deepest tha...
متن کاملEmbodied health: a guiding perspective for research in health psychology
Health psychology is based on a biopsychosocial model, which conceptualises health as a product of biological, psychological and environmental factors. However, many studies in health psychology reflect a more narrow focus on the psychosocial influences of health, leaving out physical influences, while the majority of medical studies neglect the influence of psychosocial variables. We suggest t...
متن کاملEmbodied Cognition and Simulative Mindreading
Can an embodied approach to social cognition accommodate mindreading, our ability to attribute mental states to another person? Prima facie it might not. Mindreading has been conceived in terms of what Susan Hurley calls the classical sandwich picture of the mind. On this view, perception corresponds to input from world to mind, action to output from mind to world, and cognition as sandwiched i...
متن کاملEmbodied cognitive science
The paper provides an introduction to the field of embodied cognitive science from a biological and behavioural perspective. We show how the field of neuro-ethology can help transform cognitive science from a representational to an embodied perspective. The transformation is necessary to introduce a bottom-up approach to understanding cognition in order to resolve some fundamental problems with...
متن کاملذخیره در منابع من
با ذخیره ی این منبع در منابع من، دسترسی به آن را برای استفاده های بعدی آسان تر کنید
عنوان ژورنال:
- Artif. Intell.
دوره 149 شماره
صفحات -
تاریخ انتشار 2003